Papers with error analysis of
The Gutenberg Dialogue Dataset (2021.eacl-main)
Copied to clipboard
| Challenge: | Current open-domain dialogue datasets offer a trade-off between quality and size . we build a dataset of 14.8M utterances in English and smaller datasets in german, Dutch, Spanish, Portuguese, Italian, and Hungarian . |
| Approach: | They build a high-quality dialogue corpus of 14.8M utterances in English using public-domain books from Project Gutenberg. |
| Outcome: | The proposed datasets show that the extracted dialogues are more accurate and more accurate than the larger Opensubtitles dataset. |
Error Analysis and the Role of Morphology (2021.eacl-main)
Copied to clipboard
| Challenge: | Using morphological features does improve error prediction across tasks, but is less pronounced in morphology-complex languages. |
| Approach: | They propose to use morphological features to improve error prediction across four different tasks and up to 57 languages to test their hypothesis. |
| Outcome: | The proposed model is more discriminative in morphologically simple languages than in simple ones. |
Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments (2024.lrec-main)
Copied to clipboard
| Challenge: | sarcasm detection remains an issue for both humans and natural language processing models . |
| Approach: | They analysed 300 comments from the FigLang 2020 Reddit Dataset and 39 non-native speakers of English to see if they were sarcastic. |
| Outcome: | The results show that the models and models have similar performance and weaknesses when the comments include political topics or are phrased as questions. |